Anthropic2026-08-16 07:57:52Anthropic raises misalignment risk rating and reveals unreleased Model 2Anthropic said in a company-wide risk report published on Aug. 14 that it raised its rating for catastrophic harm caused by misalignment in high-risk scenarios from “very low” to “low” under its Responsible Scaling Policy, RSP v3.4. The company said the change was not triggered by a model failing safety tests. Instead, it pointed to recent cybersecurity evaluation disclosures that increased uncertainty, along with a more technical issue: the internal benchmark it uses to detect whether models have crossed the most dangerous capability threshold has become saturated, meaning scores have effectively hit the ceiling and can no longer measure incremental gains in capability. The same report also disclosed an internal system called Model 2 for the first time. Anthropic said the model is slightly more capable than its frontier model Mythos 5 and is already used extensively inside the company. At the same time, it has not yet completed the full set of pre-deployment evaluations that would normally be required before release, and there are currently no plans to launch it externally. The disclosure comes as Anthropic is pushing toward an IPO, putting fresh attention on the gap between frontier-model capability and the tools used to evaluate safety.1010
Anthropic2026-08-15 01:55:08Anthropic says its internal Model 2 is stronger than Mythos 5 and has no release planAnthropic has used its second Risk Report to confirm, for the first time, that it is running an internal model called Model 2 that outperforms Mythos 5. The report covers risk assessments through July 15, 2026, and says the company does not currently plan to release the model publicly. Anthropic said Model 2 showed a “noticeable improvement” on internal tasks and, along with Mythos 5, has been used heavily for coding, agent work, and data generation. In AECI, Model 2 scored 162.79 versus 161.29 for Mythos 5 and 158.91 for Mythos Preview. On CoBench, which measures performance on Anthropic’s real research and engineering tasks, Model 2 posted 62.8%, compared with an 85% success rate for Anthropic’s human researchers. The report also raised the company’s misalignment risk rating in high-risk settings from “very low” to “low.” Anthropic said it reviewed more than 140,000 evaluation records in late July and found Claude had breached three real companies during cybersecurity testing. The company also disclosed five security-process failures. At the same time, Anthropic said some of its most specific task-based evaluations have become “saturated,” making further capability gains harder to measure. The disclosure arrives as OpenAI reportedly delays Astra over unresolved cyberattack concerns, setting up a contrast in how the two companies are handling frontier systems they do not plan to release.1520